Back

Journal of Chemical Theory and Computation

American Chemical Society (ACS)

Preprints posted in the last 90 days, ranked by how well they match Journal of Chemical Theory and Computation's content profile, based on 140 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.

1
Light Martini water accelerates sampling in coarse-grained molecular dynamics simulations

Elgendy, A.; Zeipelt, A. P.; Schäfer, L. V.

2026-08-03 biophysics 10.64898/2026.08.03.741232 medRxiv
Top 0.1%
57.1%
Show abstract

Molecular dynamics (MD) simulations of slow biomolecular processes, such as exploration of the conformational ensembles of intrinsically disordered proteins (IDPs), are computationally demanding. Although coarse-grained (CG) models can substantially speed up the simulations compared to all-atom MD, the sampling challenge can still be significant for large systems and long time scales. Here, we present light Martini water, a low-viscosity water model that accelerates sampling in MD simulations with the Martini CG force field. We systematically reduced the mass of the Martini water beads and verified stable, accurate integration of the equations of motion with 20 fs time steps, as typically used in Martini simulations. Light Martini water has a reduced mass of 20 amu (compared to 72 amu in the standard water model), yielding up to a 2.68-fold increase in the sampling rate of IDP chain reconfiguration in water and a 16 % increase in the lateral diffusion of lipids in a POPC bilayer. Equilibrium properties remained unaffected by the mass scaling, and the speedup was achieved without compromising simulation accuracy. The water model is trivial to implement, has no computational overhead, and should be universally applicable to Martini simulations.

2
Toward Robust Characterization of Dynamic Binding Pockets: Lessons from the HBV Capsid Assembly Modulator Site

Perez-Segura, C.; Scott, L. W.; Zlotnick, A.; Hadden-Perilla, J. A.

2026-08-10 biophysics 10.64898/2026.08.06.743403 medRxiv
Top 0.1%
40.6%
Show abstract

Protein function often depends on ligand binding pockets that fluctuate among conformational states, altering their size, shape, topology, and accessibility, yet quantitative comparison of these dynamic cavities remains challenging because their boundaries are often inherently ambiguous. The measure volinterior algorithm uses fuzzy-boundary detection to characterize enclosed molecular spaces; here, the hepatitis B virus (HBV) capsid assembly modulator (CAM) binding site is used as a model system to develop and validate a practical workflow for applying the method to dynamic protein binding pockets. The resulting methodology provides practical guidance for parameter selection and evaluation, establishes a standardized protocol for quantitative characterization of the HBV CAM pocket, and demonstrates robust, reproducible performance across conformational ensembles derived from molecular dynamics (MD) simulations. More broadly, this work provides a reproducible strategy for adapting measure volinterior to other dynamic binding pockets, enabling consistent comparison of pocket geometry among independent structural studies. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/743403v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@d460b5org.highwire.dtl.DTLVardef@11915c2org.highwire.dtl.DTLVardef@1e37524org.highwire.dtl.DTLVardef@1fc1b7_HPS_FORMAT_FIGEXP M_FIG C_FIG

3
A two-bead-per-aminoacid coarse-grained MD model with hydrogen bonding (2BPA-HB) to probe DNAJB6b-mediated suppression of polyglutamine aggregation in Huntingtons disease

ADUPA, V.; Polet, J. D.; Dekker, M.; Onck, P. R.

2026-08-27 biophysics 10.64898/2026.08.24.746793 medRxiv
Top 0.1%
40.5%
Show abstract

Polyglutamine (polyQ) aggregation plays a central role in several neurodegenerative diseases, including Huntington's disease. DNAJB6b, a molecular chaperone involved in protein quality control, is known to efficiently suppress polyQ aggregation, but its anti-aggregation mechanism remains unclear. In this work we investigate the interaction between DNAJB6b and the polyQ region (Q48) of mutant Huntingtin Exon 1 (mHttEx1) using a custom-built coarse-grained molecular dynamics model. The model incorporates a two-bead-per-amino-acid representation with hydrogen bonding (termed 2BPA-HB), and is calibrated against all-atom molecular dynamics data in terms of geometry, hydrophobicity, and hydrogen bonding. The model reproduces the tertiary structure of DNAJB6b and its interactions with Q48, and reveals an inverse correlation between DNAJB6b concentration and Q48 aggregation propensity. Our simulations show that DNAJB6b co-condensates with polyQ molecules, thereby shielding the polyQ from forming the intermolecular hydrogen bonds necessary for amyloid formation. The 2BPA-HB CGMD model en- ables efficient exploration of DNAJB6b conformations, supporting future studies of chaperone-mediated aggregation suppression and therapeutic development.

4
Ab initio side-chain sampling with PUD+ enables high-fidelity protein dynamics across AI-driven and classical simulations

Wu, D.; Wang, T.

2026-08-11 biophysics 10.64898/2026.08.10.743906 medRxiv
Top 0.1%
39.3%
Show abstract

The fidelity of molecular dynamics (MD) simulations fundamentally depends on the quality and coverage of the ab initio data used to parameterize the underlying force field, yet the role of side-chain conformational space remains insufficiently explored. In this study, we systematically investigate how comprehensive ab initio sampling of dipeptide conformations--specifically targeting side-chain degrees of freedom--impacts force field accuracy and MD simulation predictive power. We present the Protein Unit Dataset Plus (PUD+), a 40-million-conformation quantum mechanical dataset featuring unprecedented coverage of both backbone and side-chain conformational space. Machine learning force fields trained on PUD+ and integrated into AI2BMD simulations demonstrate superior energy and force prediction accuracy, capturing high-fidelity protein folding dynamics and the conformational flexibility of long-side-chain systems. Furthermore, leveraging PUD+ to reparameterize the CMAP term of the classical ff19SB force field markedly improves the description of intrinsically disordered protein (IDP) dynamics and IDP-ligand binding. Collectively, these results demonstrate that ab initio sampling of dipeptide side-chain conformations enables high-fidelity modeling of protein dynamics across both AI-driven and classical simulation paradigms.

5
pHaseMD4AI: Phase-Space Dynamics Dataset with Chemical and pH Perturbations for Physically and Kinetically Consistent Biomolecular AI

Song, T.; Guo, Y.; He, J.; Liu, Z.; Low, M.; Wang, K.; Zhang, Y.; Li, Z.; Huang, Y.; Wang, Y.

2026-07-23 bioinformatics 10.64898/2026.07.20.739474 medRxiv
Top 0.1%
39.1%
Show abstract

Protein function emerges from dynamic conformational ensembles and transitions that are challenging to characterize experimentally and computationally. Recent advances in generative AI have created new opportunities for learning molecular thermodynamics, kinetics, and conformational evolution directly from simulation data, but progress is limited by the availability of large-scale datasets that combine rigorous sampling, complete phase-space information, and diverse physicochemical perturbations. Here, we present pHaseMD4AI, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins. The peptide branch includes a complete set of canonical tripeptide and tetrapeptide systems together with post-translationally modified (PTM) and protonation-state datasets, providing synchronized atomic coordinates (R), velocities (V), forces (F), and Markov state model-based kinetic annotations. An accompanying web portal (https://isb.zju.edu.cn/md4ai/) enables users to browse, visualize, and download trajectories, annotations, and metadata. As an example application, we demonstrate a sequence-based model that can predict residue-level equilibrium dihedral distributions from sequence. pHaseMD4AI provides a resource for developing and benchmarking molecular machine learning methods while supporting broader studies of biomolecular dynamics under sequence, post-translational modification, and protonation-state perturbations.

6
Towards transferable explicit-solvent coarse-grained models for biomolecular condensates

Toplek, F. B.; Borges-Araujo, L.; Lindorff-Larsen, K.; Everaers, R.; Souza, P. C. T.; Morozova, T. I.

2026-08-29 biophysics 10.64898/2026.08.27.747511 medRxiv
Top 0.1%
38.4%
Show abstract

Biomolecular condensates formed by intrinsically disordered proteins require molecular models that accurately describe proteins in both dilute solution and condensed phases. Explicit-solvent coarse-grained models offer an attractive balance between chemical resolution and computational efficiency. Yet, it remains unclear whether improving dilute-state properties is sufficient to obtain an accurate description of condensates. Here, we address this question by introducing minimal modifications to the Martini 3 force field that combine recent advances in bonded interactions with refined protein-water interactions and strengthened glycine self-interactions, while preserving the underlying chemical transferability of the model. The resulting model substantially improves the description of single-chain conformations across a diverse benchmark of disordered proteins. We then investigate phase separation of the well-characterized low-complexity domain of heterogeneous nuclear ribonucleoprotein A1 and its sequence variants. The model reproduces several key physicochemical properties of biomolecular condensates, including chain expansion in the dense phase, sequence-dependent intermolecular contacts, protein diffusion and its relation to single-chain dimensions, and hydration, while revealing quantitative limitations in condensate density, phase equilibria, and ion partitioning. Our results show that improving dilute-state behaviour translates into a better description of condensed-phase properties, including condensate density, but is not sufficient to quantitatively reproduce the equilibrium between the dilute and dense phases.

7
Coarse-grained models for simulations of double-stranded nucleic acids for mixed protein-nucleic acid condensates

Yasuda, I.; Tesei, G.; Yamamoto, E.; Yasuoka, K.; Lindorff-Larsen, K.

2026-08-20 biophysics 10.64898/2026.08.14.744942 medRxiv
Top 0.1%
38.3%
Show abstract

Biomolecular condensates function as membraneless compartments, and some protein condensates can selectively concentrate single-stranded nucleic acids while excluding double-stranded nucleic acids. Understanding how nucleic acid structure affects partitioning into condensates has important implications for nucleic acid activity and function within condensates. Here, we present a set of coarse-grained two-bead-per-nucleotide models for simulations of double-stranded RNA and DNA in the CALVADOS framework. Our models separately represent the backbone and base, and maintain the helical structures using an elastic network potential tuned to capture chain stiffness. For dsRNA, the base stickiness was tuned using experimental data on differential partitioning of single- and double-stranded RNA into Ddx4N1 condensates in order to account for reduced base accessibility upon duplex formation. This RNA structural selectivity varied with the balance of electrostatic and non-electrostatic interactions, as revealed by simulations of condensates of the CAPRIN1 disordered region at varying ionic concentrations and with an R-to-K sequence variant. Finally, we developed parameters for double-stranded DNA using a similar approach. We envision that the CALVADOS models for double-stranded RNA and DNA will be useful for studying co-condensates of proteins and structured nucleic acids.

8
Development of force-field corrections for the RNA A-bulge motif

Kudo, T.; Ekimoto, T.; Yamane, T.; Ikeguchi, M.

2026-08-27 biophysics 10.64898/2026.08.26.747445 medRxiv
Top 0.1%
38.3%
Show abstract

Many functional RNA motifs adopt structures that deviate from the canonical A-form helix and are emerging targets for RNA-directed therapeutics. The microtubule-associated protein tau (MAPT) A-bulge motif (5'-GCAGU/5'-ACGU) is one such motif. Because its structure is stabilized by a delicate balance of local interactions, its accurate modeling remains a major challenge for molecular dynamics (MD) simulations. The experimentally determined nuclear magnetic resonance (NMR) structure of the MAPT A-bulge motif provides a stringent test of whether RNA force fields can accurately reproduce the experimentally observed conformation. Most current AMBER-family RNA force-field models have incorrectly favored a non-native base-triple state of the MAPT A-bulge motif over the experimentally observed stacked state. Structural comparison of the stacked and base-triple conformations revealed that overly favorable NH-N hydrogen bonds between the bulged adenosine and an adjacent Watson-Crick base pair were the primary source of this imbalance. We developed gHBfix-18Ab, an 18-component hydrogen-bond correction that distinguishes NH and NH2; donors. gHBfix-18Ab was combined with the previously developed OL3CP and NBfix0BPh corrections to generate the composite model gHBfix-18Ab*. This model restored the experimentally observed stacked state as the global minimum in the calculated free-energy profile and improved agreement with NMR-derived distance data for the A-bulge region. Importantly, gHBfix-18Ab* did not produce marked structural destabilization of the cUUCGg tetraloop, a widely used benchmark for RNA force-field validation, suggesting that the refinement preserves the stability of the unrelated RNA motif. These results demonstrate that targeted refinement of hydrogen-bond interactions provides a practical strategy for systematic improvement of RNA force fields toward more accurate modeling of noncanonical RNA motifs.

9
Coarse-grained simulations of long intrinsically disordered proteins: a benchmark of Martini 3 force-fields

Goss, C.; Aponte-Santamaria, C.; Gräter, F.

2026-07-17 biophysics 10.64898/2026.07.17.739185 medRxiv
Top 0.1%
38.1%
Show abstract

Martini 3 is a force field ideally suited to simulating long intrinsically disordered proteins (IDPs) in cell-like surroundings. So far, most Martini 3 variations intended for IDPs have only been benchmarked on shorter IDPs of up to 140 amino acids. In this paper, we present a comprehensive benchmark including IDPs up to 809 amino acids in length and compare the behavior of four well-known Martini 3 variations for IDPs. Modifications to only the bonded parameters result in excessively compact conformations, thereby failing to reproduce the experimental radius of gyration observed for large IDPs. In contrast, general rescaling of interaction parameters, including tuning electrostatic interactions in the case of highly-charged long IDPs, yields acceptable levels of compaction at all tested length scales.

10
DNA Sequence and Histone Variant H2A.Z Jointly Govern Nucleosome Unwrapping Pathways

GHOSH MOULICK, A.; Patel, R.; Rajpersaud, T.; Loverde, S.

2026-06-18 biophysics 10.64898/2026.06.15.732425 medRxiv
Top 0.1%
37.2%
Show abstract

Nucleosome unwrapping governs chromatin accessibility and gene regulation, yet the molecular determinants of unwrapping directionality remain poorly understood. Using atomistic and SIRAH coarse-grained umbrella sampling simulations, we show that DNA sequence and histone variant composition jointly tune a directional preference for nucleosome unwrapping. For both the ASP and Widom-601 sequences, unwrapping initiates asymmetrically from a preferred DNA end, with progressive disengagement of the H3 N-terminal tail providing the molecular switch that determines directionality in the Widom-601 system. Substitution of canonical H2A with the variant H2A.Z reverses this directional preference, shifting unwrapping to the opposite DNA end and altering the free energy landscape. SIRAH coarse-grained simulations faithfully reproduce these sequence- and variant-dependent unwrapping pathways and their qualitative free energy features, though quantitative barrier heights differ from atomistic values, identifying a target for further force field refinement. Together, these results establish H3 tail - DNA disengagement as a key mechanistic determinant of unwrapping directionality, reveal how a single histone variant substitution can reverse this preference, and validate SIRAH as an efficient framework for large-scale chromatin simulations.

11
Gaussian Accelerated Molecular Dynamics in GROMACS

Yang, Y.

2026-08-10 biochemistry 10.64898/2026.08.10.743837 medRxiv
Top 0.1%
35.4%
Show abstract

Gaussian accelerated molecular dynamics (GaMD) enhances conformational sampling by adding a smooth boost potential without requiring predefined collective variables, but an engine-integrated implementation has not been available in GROMACS. Here, we implement total-, dihedral-, and dual-boost GaMD in GROMACS 2025.4, including staged energy-statistics collection, GPU-based bias evaluation and force scaling, restart support, and outputs required for cumulant-based free-energy reweighting. The implementation was evaluated using four benchmark systems spanning conformational free energies, protein folding, and ligand recognition. For alanine dipeptide, a reweighted 100 ns GaMD trajectory recovered the major free-energy basins and rotational barriers in overall agreement with a 1000 ns conventional MD simulation. For chignolin and TC5b, all three independent trajectories for each system sampled native-like folded states from extended conformations within 300 ns and 1 s, respectively; the best TC5b structure had a minimum backbone RMSD of 0.03 nm from the experimental structure. In the benzene-T4 lysozyme system, two of five independent 500 ns trajectories captured both ligand binding and dissociation, yielding a bound pose with a minimum ligand RMSD of 0.06 nm from the crystal structure. Across all four systems, the boost-potential distributions were approximately Gaussian, and second-order cumulant reweighting resolved the expected conformational and binding free-energy basins. These results demonstrate that GROMACS-GaMD provides a practical, GPU-enabled, collective-variable-free enhanced-sampling framework for biomolecular free-energy calculations, protein folding, and ligand-binding studies.

12
Collinearity of Decomposed Energy Terms in MM-GBSA Binding Free Energy Calculations

Sevim, A.; Kocak, A.

2026-06-29 biophysics 10.64898/2026.06.24.734195 medRxiv
Top 0.1%
30.8%
Show abstract

The molecular mechanics-generalized Born surface area method (MMGBSA) is one of the most commonly used end state approaches used for the calculation of the binding free energy towards computational drug design and screening studies. It is customary to break up the free energy into van der Waals, electrostatic, polar solvation (GB), and nonpolar solvation (SA) terms and then either correlate these terms with experiment or assign physical meaning to each term. Here, we demonstrate that this assumption of independent fitting coefficients for decomposed energy terms could be invalid. Through analytic derivation and large-scale molecular dynamics simulations, we show that (i) the protein and ligand Coulomb interaction energy and the GB solvation correction are almost perfectly collinear (R2[≥]0.99) reflecting their designed role as vacuum electrostatics plus solvent screening, and (ii) the van der Waals interaction and SA term likewise exhibit strong correlation, as both depend primarily on buried surface area. Interaction entropy and C2 entropy corrections are also found to be strongly dependent on underlying electrostatic fluctuations, further reinforcing redundancy. These findings hold both at the level of instantaneous trajectory fluctuations and when averaged across a diverse set of 139 protein-protein complexes and persist in both single-trajectory and three trajectory MMGBSA protocols. Our results caution against using decomposed MMGBSA terms as independent predictors in regression models and suggest instead combining correlated terms into effective polar, nonpolar, and entropic contributions. Our study provides a systematic diagnosis of collinearity in MMGBSA and highlights pathways toward more interpretable and statistically robust predictive modeling.

13
Validation of MHPC512: A Publicly Available Special-Purpose Supercomputer for Molecular Dynamics Simulations of Biomolecular Systems

Chen, L.;Chen, Y.;Cheng, Z.;Guo, J.;He, M.;Li, H.;Li, X.;Li, Z.;Ma, J.;Ma, S.;Peng, C.;Qian, C.;Qu, Z.;Sun, X.;Tang, X.;Wang, Y.;Yu, B.;Zhai, Y.;Zhang, B.;Zhang, S.;Zhang, S.;Hu, Z.;Shan, Y.;Mei, Y.

2026-06-11 Molecular Biology 10.64898/2026.06.10.727820 medRxiv
Top 0.1%
27.3%
Show abstract

MHPC512 is a massively parallel, special-purpose supercomputer designed primarily for atomic-level molecular dynamics (MD) simulations of biomolecular systems. It comprises 512 processor units interconnected by a high-speed three-dimensional torus network and employs a custom chip architecture that uses 35-bit fixed-point arithmetic to accelerate computation while controlling precision loss within an acceptable margin. The preprocessor is compatible with GROMACS and AMBER input formats and supports widely used biomolecular force fields (including CHARMM, AMBER, and OPLS/AA), the Neutral Territory method for short-range nonbonded interactions, the k-space Gaussian Split Ewald method for long-range electrostatics, and multiple thermostats, barostats, and integrators. We present a three-tier validation protocol--comparing static energy and virial components, examining ensemble distributions (NVE, NVT, NPT), and evaluating long-time statistical properties--demonstrating that MHPC512 reproduces results consistent with GROMACS and AmberTools. Application examples, including bulk water, dipeptide conformational sampling, folding of fast-folding peptides, membrane-protein systems, lipid self-assembly, and GPCR conformational transitions, further confirm its reliability. MHPC512 has been deployed at multiple supercomputing centers and is publicly accessible, representing a significant advance in high-throughput, large-scale biomolecular MD simulations.

14
AdaptivePy: a unified Python framework for adaptive sampling in molecular dynamics

Nadeem, H.; Kleiman, D. E.; Shukla, D.

2026-08-06 biophysics 10.64898/2026.08.04.742868 medRxiv
Top 0.1%
27.2%
Show abstract

Adaptive sampling accelerates the exploration of conformational space in molecular dynamics (MD) simulations by repeatedly analyzing the accumulated trajectories and seeding a new round of simulations from informative configurations. A growing collection of adaptive sampling policies has been proposed, each built around a particular notion of what makes a configuration informative, yet these methods are scattered across separate and often incompatible implementations, which complicates their systematic comparison and their combined use in meta adaptive sampling schemes. Here, we present AdaptivePy, a compact and extensible Python framework that implements nine seed-selection policies behind a single configuration-driven interface, spanning simple population-based baselines, several established machine-learning and geometry-based methods, and two ensemble or meta sampling policies introduced in this work. We show that the shared implementation reproduces the characteristic selection behavior of each policy on a series of analytic benchmark landscapes. We also introduce a new adaptive sampling scheme that employs TS-DAR, a deep learning framework originally designed to identify transition states, into an acquisition criterion that drives the discovery of an entire multi-basin landscape starting from a single basin. We further demonstrate that the common interface enables meta adaptive sampling policies, which aggregate the rankings of several policies into a single set of seeds. AdaptivePy thereby provides a unified testbed for the adoption, benchmarking, and continued development of adaptive sampling methods for biomolecular MD simulations.

15
Enhanced-Sampling Molecular Dynamics Recovers Rare Functional RNA Conformations Across Diverse Structural Contexts

Li, D.; Ken, M.

2026-08-24 biophysics 10.64898/2026.08.22.746453 medRxiv
Top 0.1%
26.1%
Show abstract

Accurate determination of RNA conformational ensembles is essential for understanding RNA function and advancing RNA-targeted drug discovery, yet lowly-populated alternative states remain difficult to resolve with atomistic detail. A central constraint is that experimental refinement can only select conformations already present in the starting library, making library generation the limiting step. Using the HIV-1 trans-activation response element (TAR) as a model system, we benchmarked conventional MD (cMD) against the enhanced sampling methods Gaussian-accelerated MD (GaMD), replica-exchange Gaussian-accelerated MD (Rex-GaMD), replica-exchange with solute tempering (REST2), and temperature replica-exchange MD (T-REMD), as well as the structure-prediction based methods FARFAR2 and AlphaFold 3. Each library was refined against experimental residual dipolar couplings (RDC) and validated independently using ensemble-averaged QM/MM chemical shifts. We showed that T-REMD produced the most accurate ensemble by both measures, and its advantage tracked with broader, more continuous coverage of the interhelical conformational landscape. Broad temperature-range T-REMD also sampled conformations resembling excited state 1 (ES1) and the U23-A27-U38 base-triple, without requiring these states to be specified during library generation. More accurate ensembles further improved coverage of experimentally observed ligand-bound TAR conformations and enhanced ensemble-based virtual screening, linking structural accuracy to functional utility. The same workflow applied to the preQ1 class I riboswitch and the UUCG tetraloop improved agreement with experimental data in both cases. Together, these results establish replica-exchange enhanced sampling, particularly T-REMD, as an effective strategy for constructing experimentally validated RNA ensembles and accessing conformations corresponding to rare functional substates.

16
RPDynaFlow: Generating RNA-Protein Conformational Ensembles by Atomic Conditional Flow Matching

Li, Y.; Lu, K.

2026-08-28 biophysics 10.64898/2026.08.28.747734 medRxiv
Top 0.1%
25.9%
Show abstract

Conformation ensembles of biomolecules provide the basis for understanding structural transformations and drug design. Deep-learning generative models have advanced protein and small molecule ensemble generation, while RNA-Protein complexes remain unaddressed due to the chemical heterogeneity, limited dataset size and the different flexibility scales of RNA and protein components. We present RPDynaFlow, a flow-matching model to generate conformation ensembles of RNA-protein complexes, trained on 600 ns trajectories of molecular dynamics(MD) simulation. The results show our model extends the sampling range of the phase space compared to MD simulation, which couldbe treated as a rapid and efficient complement to MD trajectoriesfor studying RNA-protein interactions.

17
MOFF2: A Transferable Coarse-Grained Protein Force Field for Predictive Condensate Simulations

Liu, S.; Zhang, Y.; Riveros, I.; Wang, C.; Zhang, B.

2026-06-10 biophysics 10.64898/2026.06.10.731384 medRxiv
Top 0.1%
25.6%
Show abstract

Coarse-grained protein force fields enable simulations of biomolecular systems at length and time scales that are difficult to access with atomistic models, but achieving transferability across folded, intrinsically disordered, and multidomain proteins remains challenging. A central difficulty is that one-bead-per-residue models must represent chemically specific residue interactions while also absorbing solvent-mediated and many-body effects into a simplified energy function. Here, we present MOFF2, a transferable coarse-grained protein force field that combines residue-pair-specific interactions with a density-dependent many-body potential. MOFF2 is optimized using a two-stage strategy: bottom-up parameter learning from heterogeneous reference ensembles followed by refinement against experimental conformational observables. The resulting model provides balanced performance across ordered proteins, intrinsically disordered proteins, and multidomain proteins, and predicts condensate saturation-concentration trends for A1-LCD variant systems. Analysis of the learned parameters reveals chemically interpretable interaction patterns and density-dependent effects that explain the models improved transferability. These results demonstrate that combining a generalized coarse-grained energy function with data-driven optimization can produce a practical and interpretable force field for protein conformational and condensate simulations.

18
Sequence-dependent conformational and mechanical landscapes of double-stranded nucleic acids

Sharma, R.; Patelli, A. S.; Singh, R.; Petkeviciute-Gerlach, D.; Gonzalez, O.; Maddocks, J. H.

2026-08-10 biophysics 10.64898/2026.08.09.740023 medRxiv
Top 0.1%
23.0%
Show abstract

The sequence-dependent mechanical landscapes of double-stranded nucleic acid (dsNA) remain largely unexplored beyond canonical dsDNA. We describe cgNA+, a coarse-grained predictive model of the mechanics of dsRNA, DNA:RNA hybrids, and epigenetically modified dsDNA, all parameterised from 1.26 milliseconds of atomistic simulations. cgNA+ predicts non-local sequence-dependent equilibrium shape and stiffness with errors an order of magnitude smaller than sequence-variability, while enabling exploration of numbers of sequences inaccessible to atomistic simulation. We show that dsNA equilibrium shape is strongly influenced by flanking sequence up to octamer context, with flexible dimer-steps more context-sensitive. CpG-modification alters equilibrium shape comparable to changes caused by single-nucleotide polymorphisms. Groove width analysis across dsNA decamers reveals strong sequence dependence, reflecting the differing characteristic helical geometry of dsDNA and dsRNA, whereas DRHs exhibit mixed behaviour depending on DNA-strand pyrimidine content. CTCF binding sites exhibit a distinct groove width signature. Persistence-length spectra from [~] 9 million sequences indicate that dsRNA is stiffer than dsDNA, whereas DRH exhibit intermediate stiffness modulated by DNA strand pyrimidine content. Persistence length increases upon CpG-modification, but decreases on hypermodification. Overall, the cgNA+ model enables a first, highly accurate, very large-scale, comparative study of sequence-dependent mechanics both within and across dsNA classes, demonstrating previously hidden regulatory layers. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=74 SRC="FIGDIR/small/740023v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@6de6acorg.highwire.dtl.DTLVardef@1435181org.highwire.dtl.DTLVardef@9c2b60org.highwire.dtl.DTLVardef@e3abf0_HPS_FORMAT_FIGEXP M_FIG C_FIG

19
Parallel DNA Holliday Junctions: Myth or Reality? An Atomistic Molecular Dynamics Study

Lemmens, T.; Sponer, J.; Stadlbauer, P.; Krepl, M.

2026-07-16 biophysics 10.64898/2026.07.10.737670 medRxiv
Top 0.1%
22.7%
Show abstract

Holliday junctions (HJs) are key intermediates of homologous recombination and fundamental building blocks in DNA nanotechnology. Although the canonical antiparallel stacked-X conformation has been extensively characterized by X-ray crystallography, whether parallel HJ conformations exist in aqueous solution remains unresolved. Here, we address this question using extensive atomistic molecular dynamics (MD) simulations and replica-exchange umbrella sampling (REUS) free-energy calculations. Starting from canonical antiparallel HJs, standard MD simulations occasionally revealed spontaneous transitions to parallel conformations on the microsecond timescale without disrupting the DNA duplexes or passing through an open junction intermediate. REUS free-energy profiles confirmed the antiparallel state as the global minimum but also identified the parallel conformation as a well-defined local minimum, indicating it is thermodynamically metastable despite an estimated solution population below 1%. The combination of low equilibrium occupancy, microsecond interconversion dynamics, and the surprisingly close structural similarity between antiparallel and parallel junctions provides a plausible explanation for the lack of direct experimental detection. We further found that the free-energy landscape is only weakly affected by branching-point sequence and salt concentration. These results reconcile the apparent absence of experimental evidence for parallel HJs with their structural feasibility in solution and offer a fresh perspective on the historical debate. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=61 SRC="FIGDIR/small/737670v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@bb61a0org.highwire.dtl.DTLVardef@673a7org.highwire.dtl.DTLVardef@192f5adorg.highwire.dtl.DTLVardef@13f4941_HPS_FORMAT_FIGEXP M_FIG C_FIG

20
Predictive all-atom simulations of disordered proteins and biomolecular condensates through osmometry-guided force-field optimization

Ivanovic, M. T.; von Roten, V.; Schuler, B.; Best, R. B.

2026-08-26 biophysics 10.64898/2026.08.25.747127 medRxiv
Top 0.1%
22.4%
Show abstract

All-atom simulations with explicit solvent provide the most detailed and accurate description of dynamics and mechanisms in intrinsically disordered proteins and their condensates. However, interactions involving charged residues and ions remain a persistent source of systematic error. Here we introduce an osmometry-guided optimization strategy that directly targets residue-residue, residue-ion and ion-ion interactions. Osmotic pressure provides key experimental information on molecular interactions and can be calculated directly and rapidly from simulations, enabling efficient iterative force-field optimization. The resulting parameters improve agreement of all-atom simulations with a range of experimental data: single-molecule FRET measurements for 16 monomeric intrinsically disordered regions; NMR relaxation data for a complex between an IDP and a folded protein domain; and mean FRET efficiencies and chain reconfiguration times of IDPs in biomolecular condensates of highly charged proteins. For such condensates, simulations with an osmometry-calibrated force field provide the missing link for predicting condensate dynamics across length and time scales. The presented optimization strategy is broadly extensible to other interaction classes, including those governing protein-DNA and protein-RNA assemblies.